Skip to content

Chunk recipe execution and output processing - #1017

Open
ValterH wants to merge 1 commit into
mem-opt/base-recipe-executionfrom
mem-opt/6-chunk-recipe-execution
Open

ValterH wants to merge 1 commit into
mem-opt/base-recipe-executionfrom
mem-opt/6-chunk-recipe-execution

Conversation

@ValterH

@ValterH ValterH commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Split from #996 to keep each review focused on one behavior or optimization.

Transform queries and post-process member outputs in passes over rows within the chunk memory limit in RecipeExecution.

Keep member outputs in the model output dtype until post-processing, which casts each pass to the callback-prepared query dtype and inverts numerical targets there, rather than in ICLModel.forward and predict. The benchmark adapter releases transformed queries before post-processing.

Document the row independence required of fitted feature and output steps and inverse target steps in Recipe.

Depends on #1016 and #1009. The dependency-only base includes #1013, #1014, #1016 and the rebased #1009 stack to isolate this commit for review; rebase and retarget to main after those dependencies merge. Split from #996; related to #994.

@copy-pr-bot

copy-pr-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@RBendias
RBendias marked this pull request as ready for review September 28, 2026 21:49
- Keep member outputs in the model output dtype until post-processing,
  which casts them to the callback-prepared query dtype and inverts
  numerical targets.
- Transform queries and post-process member outputs in passes over rows
  within the chunk memory limit in `RecipeExecution`.
- Require equal row counts across ensemble groups only when passes split
  them.
- Post-process the benchmark adapter's outputs through `transform_output`
  after freeing the transformed queries.

Signed-off-by: Jingang Qu <jqu@nvidia.com>
Co-authored-by: Cedric Lorenz <clorenz@nvidia.com>
@ValterH
ValterH force-pushed the mem-opt/6-chunk-recipe-execution branch from 6a42da5 to a4a75ea Compare September 29, 2026 07:44
@ValterH
ValterH force-pushed the mem-opt/base-recipe-execution branch from 82e1773 to ff8ffdd Compare September 29, 2026 07:44

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants